Papers with Evaluation metrics

4 papers
MapQaTor: An Extensible Framework for Efficient Annotation of Map-Based QA Datasets (2025.acl-demo)

Copied to clipboard

Challenge: Mapping and navigation services struggle to handle natural language geospatial queries.
Approach: They introduce an extensible open-source framework that streamlines the creation of reproducible, traceable map-based QA datasets.
Outcome: a new open-source framework streamlines the creation of reproducible, traceable map-based QA datasets.
SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods for Open Domain Event Detection (ODED) lack representative representations of the real world, making it difficult to accurately reflect performance of various ODED methods in real-world scenarios.
Approach: They propose a scalable and reliable Semantic-level Evaluation framework for Open domain event detection by constructing a more representative evaluation benchmark and introducing a semantic evaluation metric.
Outcome: The proposed framework first constructs a more representative evaluation benchmark that currently includes 564 event types covering 7 major domains, with a cost-effective supplementary annotation strategy to ensure the benchmark’s representativeness.
Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors (2021.emnlp-main)

Copied to clipboard

Challenge: Evaluation metrics are a key ingredient for progress of text generation systems . a class of novel evaluation metrics based on BERT and its variants has been explored .
Approach: They propose to disentangle BERT-based evaluation metrics along linguistic factors . they show they are sensitive to lexical overlap, just like BLEU and ROUGE .
Outcome: The proposed metrics capture all aspects but are sensitive to lexical overlap, just like BLEU and ROUGE, the authors show .
IM^2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Evaluation metrics for dialogue systems are expensive and time-consuming . current evaluation metrics focus on a single quality or several qualities .
Approach: They propose an interpretable, multi-faceted, and controllable framework to combine dialogue metrics which are good at measuring different qualities.
Outcome: The proposed framework integrates a large number of evaluation metrics to improve the performance of the model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations